Papers with error analysis of

3 papers
The Gutenberg Dialogue Dataset (2021.eacl-main)

Copied to clipboard

Challenge: Current open-domain dialogue datasets offer a trade-off between quality and size . we build a dataset of 14.8M utterances in English and smaller datasets in german, Dutch, Spanish, Portuguese, Italian, and Hungarian .
Approach: They build a high-quality dialogue corpus of 14.8M utterances in English using public-domain books from Project Gutenberg.
Outcome: The proposed datasets show that the extracted dialogues are more accurate and more accurate than the larger Opensubtitles dataset.
Error Analysis and the Role of Morphology (2021.eacl-main)

Copied to clipboard

Challenge: Using morphological features does improve error prediction across tasks, but is less pronounced in morphology-complex languages.
Approach: They propose to use morphological features to improve error prediction across four different tasks and up to 57 languages to test their hypothesis.
Outcome: The proposed model is more discriminative in morphologically simple languages than in simple ones.
Error Analysis of NLP Models and Non-Native Speakers of English Identifying Sarcasm in Reddit Comments (2024.lrec-main)

Copied to clipboard

Challenge: sarcasm detection remains an issue for both humans and natural language processing models .
Approach: They analysed 300 comments from the FigLang 2020 Reddit Dataset and 39 non-native speakers of English to see if they were sarcastic.
Outcome: The results show that the models and models have similar performance and weaknesses when the comments include political topics or are phrased as questions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations